Back

Computers in Biology and Medicine

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Computers in Biology and Medicine's content profile, based on 128 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.

1
AutopsyPrint: A novel tool for translating ballistic and sharp force injury trajectory findings into 3D printable models

Parsons, C. E.; Thomsen, A. H.; Petersen, M. V.

2026-07-13 forensic medicine 10.64898/2026.07.10.26357738 medRxiv
Top 0.1%
20.0%
Show abstract

When autopsy findings are presented on two-dimensional paper-based models, there is an inherent reduction of spatial information, which the viewer must infer from simplified anatomical and geometrical representations. Multiple diagrams representing trajectory angles must be integrated into a complete mental model, introducing potential for errors in viewer understanding. 3D models can address these issues but have shown limited adoption in forensic autopsy reporting given the technical competences and software required to produce them. Here, we present AutopsyPrint, a workflow and open web-based tool for generating 3D printable body models annotated with wound trajectories. The tool supports marking different wound types, including ballistic and stab wounds, on male and female bodies, which can be posed to accommodate a diversity of trajectories. Based on our testing, we present a set of suggested workflow steps and parameters based to facilitate standardization of the 3D models produced, balancing between precision and print time and materials. To ensure accessibility, the tool runs fully in the user's browser, and all annotated data is stored locally. By making AutopsyPrint open access, we intend to build practical experience with model creation, to ultimately advance the use of 3D models in the field.

2
Comparing Machine Learning Approaches for Predicting CFD-Derived Stroke Risk Indicators in Atrial Fibrillation Patients

Melidoro, P.; Cavarra, R.; Mostafa, S.; Lip, G. Y. H.; Klis, M.; Williams, S. E.; Aslanidi, O.; De Vecchi, A.

2026-06-08 bioengineering 10.64898/2026.06.04.730070 medRxiv
Top 0.1%
19.2%
Show abstract

Non-valvular atrial fibrillation (AF) is associated with a five-fold increased risk of stroke, mainly due to impaired contractility of the left atrium (LA) leading to blood stasis and subsequent thrombus formation within the left atrial appendage (LAA). Current AF stroke risk stratification schemes, such as the CHA2DS2-VASc/ CHA2DS2-VA score, use comorbidities and do not capture mechanistic factors like blood flow dynamics and hypercoagulability. To address this, we developed a multiphase computational fluid dynamics (CFD) model of the LA, incorporating patient-specific geometries; modelling of the coagulation cascade; and non-Newtonian blood behaviour within the LAA. Using 84 simulation cases generated via Latin Hypercube Sampling of physiological blood parameters and 21 patient-derived LA anatomies, we trained surrogate machine learning models, including Ridge regression, XGBoost, Gaussian Process Emulators (GPEs), and deep learning networks, to predict CFD outputs such as blood viscosity in and fibrin concentrations in the LAA. Deep learning achieved R{superscript 2} values up to 0.90, with the accuracy increasing when both physiological parameters and the raw CT image were included. Other models showed uneven performance with R2 values below 0.7, highlighting the role of nonlinearities between parameters. The study presents a novel CFD model that captures the transition from blood stasis to clot formation, representing the full thrombotic continuum underlying stroke risk in AF, and a deep learning approach to enable efficient prediction of mechanistic outputs of clinical value for stroke risk stratification in AF patients. Author SummaryAtrial fibrillation is a common heart rhythm disorder that greatly increases the risk of stroke. In many patients, blood can pool inside a small pouch of the heart called the left atrial appendage, where clots may form and later travel to the brain. Current clinical tools used to estimate stroke risk mainly rely on a patients medical history and do not directly assess the mechanistic processes that lead to clot formation. In this study, we developed a computer model that simulates how blood flows and clots inside the heart using patient-specific heart anatomies derived from medical imaging. Our model combines blood flow, blood biochemistry, and the changing physical properties of blood during clot formation. We then used machine learning methods to predict these complex simulation results more efficiently. Deep learning models performed best, particularly when both clinical parameters and heart imaging data were included. Our work provides a new way to study the full process linking abnormal blood flow to clot formation in atrial fibrillation. In the future, this approach could support more personalised and mechanistic assessment of stroke risk and help guide treatment decisions.

3
Rt3DE-based finite element analysis of functional tricuspid regurgitation and RV free wall approximation

Tondi, D.; Vailetta, S.; Sturla, F.; Vismara, R.; Votta, E.

2026-07-14 bioengineering 10.64898/2026.07.13.736182 medRxiv
Top 0.1%
18.6%
Show abstract

PurposeFunctional tricuspid regurgitation (FTR) is driven by right ventricular (RV) remodeling, annular dilation, and papillary muscle dislocation. Free wall approximation (FWA) has been proposed to treat FTR by addressing RV dilation, but its effects on tricuspid valve (TV) biomechanics remain unclear. We present a real-time 3D echocardiographic (rt3DE)-based finite element framework to quantify TV biomechanics under FTR, and preliminarily apply it to assess FWA effects. MethodsSubject-specific models were developed from rt3DE data of three dilated porcine hearts in an ex-vivo mock-loop. TV geometries at end-diastole and peak systole (PS) were complemented by parametric chordae tendineae and hyperelastic tissue properties. TV closure was simulated under a standard pressure load and image-based annular motion. After tuning chordae length to replicate the PS ground truth in FTR, FWA was simulated as 30% and 60% approximations along three anatomical directions (anterior-posterior, A-P; anterior-septal, A-S; anterior-septal wall, A-SW). ResultsIn FTR simulations, median geometric errors ranged from 1.16 to 1.26 mm; median stress ranged from 56.4 to 74.7 kPa. FWA simulations predicted regurgitant orifice area (ROA) reductions by 53-99%, albeit overestimating the residual ROA vs. in vitro ground truth when starting from particularly extreme FTR conditions; concomitantly, a median stress reduction by 8-43% vs. FTR conditions was predicted. ConclusionPreliminary data suggest that our rt3DE-based framework can reliably quantify FTR-related TV biomechanics and that post-FWA biomechanics depends on initial FTR conditions. A larger cohort is required to verify the method and obtain statistically significant results.

4
The impact of the Impella RP(R) device on a failing right heart. A modelling and simulation approach to ascertain its potential

De Lazzari, B.; Richter, A.; Nix, C.; Badagliacca, R.; Pitino, A.; Gori, M.; Scoccia, G.; Capoccia, M.; DE LAZZARI, C.

2026-06-26 cardiovascular medicine 10.64898/2026.06.24.26356428 medRxiv
Top 0.1%
15.8%
Show abstract

Background and Objective: Indications for right ventricular assist device (RVAD) insertion include right heart failure after implantation of a left ventricular assist device or early graft failure following heart transplantation. This study aimed to investigate how the upstream and downstream circulatory network interacts with the Impella RP(R) device. Methods: A numerical model of the Impella RP(R) was implemented within CARDIOSIM(C) software platform for this study. In the numerical configuration, the RVAD aspirated blood from either the right atrium (RA-PA connection) or the right ventricle (RV-PA connection) and delivered it to the pulmonary artery. Only RA-PA connection is the currently used setting for Impella RP(R) in clinical practice. Based on right ventricular (RV) decompression and total flow, our study may help define the need for a direct RV-unloading Impella RP(R). Results: The simulations showed that activating the RVAD in RA-PA mode, regardless of its rotational speed, the mean pulmonary artery pressure (PAP) percentage change was higher than the unsupported condition when the mean systemic venous pressure (SVP) and the pulmonary artery wedge pressure (PAWP) were both set to 20 mmHg. When RV-PA connection was applied, a similar trend was observed although the PAP percentage changes were about halved compared to the RA-PA connection. Conclusions: The Impella RP(R) has the potential to become a valid option for RV support based on current experimental and simulation data. Although already in use, further evaluation in the clinical setting will likely confirm its potential and lead to a more routinely application for RV support.

5
Exploratory Assessment of Pulsed-Wave Doppler Representations of Lung Sounds Using Deep Learning: An In-Vitro Phantom Study

Saad, A. A.; Murthi, S. B.; Boctor, E. M.; Teeter, W. A.; Seam, N.

2026-06-10 respiratory medicine 10.64898/2026.06.09.26353787 medRxiv
Top 0.1%
15.1%
Show abstract

The increasing availability of portable ultrasound systems motivates exploration of novel approaches to respiratory signal assessment. In this in-vitro study, we investigate whether pulsed-wave (PW) Doppler ultrasound can capture structured spectral patterns from replayed lung sound recordings. Digitized respiratory sounds were replayed through a tissue-mimicking ultrasound phantom, generating 1,478 PW Doppler spectral images from recordings associated with healthy subjects and several externally labeled disease categories. Exploratory classification experiments using a ResNet-18 architecture demonstrated that these Doppler representations contain learnable differences under controlled conditions. These findings motivate further investigation into PW Doppler as a potential representation of respiratory acoustics.

6
Explainable Artificial Intelligence for Cross-Dataset Generalizable Biomarker Discovery in Cardiovascular diseases (CVDs)

Abbasi, A. F.; Sajjad, M.; Vollmer, S.; Dengel, A.; Asim, M. N.

2026-07-27 bioinformatics 10.64898/2026.07.23.740273 medRxiv
Top 0.1%
13.0%
Show abstract

CVDs are heterogeneous, multifactorial disorders that remain the leading cause of global mortality from infancy to old age. It requires an early identification and treatment of risk factors to accelerate disease prevention and morbidity improvement. Advancements in transcriptomics technologies gives large pool of heterogenous gene expression data. The technical heterogeneity of gene expression data reduces ability to compare multiple cross-platform datasets at once. To bridge gap, we systematically evaluate three data harmonization techniques: Shambhala-2, TDM, and UPC to align heterogeneous data into a shared expression space while preserving biological signals. Our pipeline integrates 25 independent datasets comprising 983 samples across 23 distinct CVDs phenotypes from both RNA-seq and microarray platforms. The framework benchmarks 35 Machine learning (ML) and Deep learning (DL) classifiers, including Transformers and ResNets, across three data modalities such as RNA-seq, microarray hybridization and RNA-seq + microarray and multiple tissue types. To ensure clinical trustworthiness, we apply multiple Explainable artificial intelligence (XAI) methods, such as SHapley additive exPlanations (SHAP) and Integrated gradientss (IGs), and assess their reliability using quantitative metrics like Area over the perturbation curve (AOPC), Sensitivity, and Infidelity. Results indicate that Shambhala-2 provides superior harmonization by maximizing the biological signal-to-platform ratio. Evaluation of XAI methods reveals that Shapley-based approaches offer the highest stability for identifying influential genomic features in high-dimensional data. Functional enrichment and pathway analyses further confirmed the involvement of identified biomarkers in key cardiovascular processes, including inflammation, immune regulation, oxidative stress, and vascular remodeling. Collectively, this study provides a scalable and interpretable road-map that integrates XAI with cross-dataset biomarker discovery, supporting the transition toward precision cardiology.

7
Next-Generation Skin Cancer Detection Using Efficient Fuzzy Fusion of Genomic and Imaging Data

Molla, A. R.; Maity, A.; Saha, S.; Bhattacharya, R.; Chakraborty, A.; Biswas, S.; Nath, S.

2026-06-08 health informatics 10.64898/2026.06.05.26355024 medRxiv
Top 0.1%
13.0%
Show abstract

Skin cancer requires early detection for improved survival rates. Most existing methods rely on deep learning based image classification, which is affected by visual similarity among lesions. Fewer studies use Gene Expression (GE) analysis, which captures molecular characteristics but lacks structural and visual details. To overcome limitations of individual modalities, this paper proposes a multimodal framework integrating dermoscopic images and GE profiles for skin cancer classification. EfficientNet and logistic regression are used for image based analysis and genomic skin lesion profiling, respectively, followed by fuzzy rule based decision systems to reduce uncertainty within individual modalities. Finally, fuzzy fusion combines predictions from both modalities using uncertainty based weighting of classifier outputs. The experimental findings show that both the image based and GE based classification models individually achieved accuracies of nearly 92%. However, the integration of prediction results through the proposed fuzzy fusion strategy further enhanced the classification performance, achieving an overall accuracy of 94.25%. The results obtained outperform contemporary methods, highlighting the effectiveness of combining complementary multimodal information compared with single modality approaches.

8
Dynamic Graph Representation Learning for Data-Driven Huntington's Disease Staging: Evaluation Against Existing Embedding Methods and State-Space Models

Abu Zohair, L. M.; Zantout, H.; Gow, A. J.; Woodward, J.; Lones, M.; Vallejo, M.

2026-06-30 health informatics 10.64898/2026.06.27.26355575 medRxiv
Top 0.1%
11.6%
Show abstract

Huntington's disease (HD) presents a heterogeneous neurodegenerative course, with motor, cognitive, and functional symptoms progressing differently across individuals. This atypical progression complicates the definition of discrete disease stages, hindering understanding of disease trajectories, timely pa- tient care, and therapy development. Consequently, current clinical staging systems rely heavily on clinician-defined, domain-specific criteria and fixed clinical measurement boundaries for stage assignment, reducing objectivity and often leading to overlapping clinical measurements across stages. While machine learning methods can help, existing approaches cannot fully capture complex temporal relationships within and across patients. We propose URL- STFN, a dynamic graph-based representation learning model that encodes both inter- and intra-patient temporal patterns from longitudinal clinical measures. We then evaluate disease stages formed through clustering and stability analysis of URL-STFN latent representations, and compare them with representations obtained from conventional embedding approaches. We further benchmark these clustering-based stages against states derived from conventional temporal models, including DHMM. We hypothesize that clustering URL-STFN latent representations enables identification of HD stages with reduced overlap in clinical measurements. The proposed framework is evaluated using 1,477 clinical visits from the Enroll-HD dataset, a large lon- gitudinal cohort with repeated clinical assessments. For staging, we used 44 clinical measurements spanning motor, cognitive, and functional domains. URL-STFN identifies clinically meaningful HD stages consistent with estab- lished disease progression while reducing overlap in clinical feature values compared with DHMM-derived and clinical staging approaches. These find- ings highlight the potential of a dynamic graph-based representation learning and clustering framework to support more objective, data-driven, and precise HD staging.

9
Comorbidity structure as an inductive bias: Comparing output-head designs for multi-label prediction of diabetes and myocardial infarction complications

Asumboya, W. A.; Agbenorhevi, P. K.; Adams, C. F.; Ayariga, D. A.; Adjadeh, T.; Adams Ziblim, S.; Kwofie, S. K.

2026-06-23 bioinformatics 10.64898/2026.06.18.733068 medRxiv
Top 0.2%
9.0%
Show abstract

BackgroundClinical complications are often predicted with separate sigmoid outputs, even when the target labels arise from related pathophysiological processes. This paper asks whether output-layer choice should reflect both predictive convenience and the biological structure assumed among complications. The central premise is that label-dependence mechanisms are explicit hypotheses about comorbidity, not generic modelling additions. MethodsOutput-head assumptions were compared across two clinically distinct multi-label prediction tasks. In Type 2 diabetes (T2D), six heads were evaluated for nephropathy, neuropathy, and retinopathy: independent baseline, linear additive, multiplicative, symmetric conditional random field (CRF), residual multilayer perceptron (MLP), and combined additive-multiplicative. In myocardial infarction (MI), four heads were evaluated for ventricular tachycardia, ventricular fibrillation, and atrioventricular block: independent baseline, linear additive, multiplicative, and symmetric CRF. All experiments used five training data fractions and seven independent seeds, with the same shared-backbone protocol within each disease setting. ResultsIn T2D, the symmetric CRF gave the most consistent improvement pattern, ranking highest at full data and at the two lowest data fractions while adding only three interaction parameters. At 20% training data, it was the only interaction head whose aggregate mean exceeded the independent baseline. The residual MLP, despite 123 interaction parameters, remained below the baseline across all T2D fractions. In MI, rankings changed across fractions: the multiplicative head led at 80% and 60%, the CRF led at 100% and 20%, and the baseline led at 40%. The combined additive-multiplicative head did not improve robustness in T2D and showed the largest negative baseline-relative deviations at lower fractions. ConclusionThe findings support a biology-guided view of output-layer design. A small constrained mechanism was most useful when its symmetry matched the shared microvascular structure of T2D, whereas the heterogeneous electrophysiology of MI produced no stable winner. Output-layer choice should therefore be reported and defended as an assumption about disease structure instead of a routine hyperparameter decision. Author summaryMany clinical prediction models treat complications as separate outcomes, even when clinicians know they often arise together. We studied whether the last layer of a model should reflect that biological knowledge. We compared several output heads across two disease settings: Type 2 diabetes, where nephropathy, neuropathy, and retinopathy share a common microvascular origin, and myocardial infarction, where electrical complications arise from a mixture of shared and location-specific mechanisms. We found that a small symmetric CRF head was most useful in the diabetes task, especially when training data were limited, while no single interaction head dominated in myocardial infarction. This suggests that modelling comorbidity is not only a technical choice; it is a statement about how disease processes relate to one another. Our results encourage researchers to report and justify output-layer design as part of the clinical modelling argument, rather than treating it as a routine hyperparameter.

10
GuavaVision AI: An Explainable Deep Learning Framework for Automated Classification, Lesion Localization, and Segmentation of Guava Diseases

Biswas, J.; Islam, M.; Bangabashi, M. M.; Akter, M.; Nishi, T. S.; Sheikh, M. K.; Mia, M. R.; Anwar, M. M.

2026-06-23 bioengineering 10.64898/2026.06.18.733093 medRxiv
Top 0.2%
8.0%
Show abstract

Guava cultivation is considerably influenced by foliar and fruit diseases whose overlapping symptoms and environmental variability make accurate field-level diagnosis challenging. Numerous studies have been conducted to find efficient methods of diagnosing plant diseases, but most focus on image-level classification and do not include lesion localization or pixel-level segmentation of the images within a single framework of analysis. This study proposes a comprehensive framework for utilizing automated image analysis to classify guava leaf and fruit diseases at the image level, locate lesions, and segment lesions at the pixel level from multiple images of the same type of disease collected from various growing conditions. The dataset was enriched through three augmentation strategies including standard preprocessing, structured augmentation, and GAN-based synthetic image generation, expanding the effective training data to approximately 7,000 images, while a 5-fold cross-validation strategy guided model selection and final performance was assessed on a held-out test set. The experimental evaluation of multiple state-of-the-art Convolutional Neural Networks (CNNs) for the classification of guava leaf and fruit diseases indicated that the model generated using the ResNet50+DenseNet121 model fusion achieved the highest classification accuracy of 98.20%. For lesion detection and segmentation, YOLOv8-seg outperformed Mask R-CNN, achieving mAP@0.5 of 0.907 and 0.889, and mAP@0.5:0.95 of 0.783 and 0.769 for detection and segmentation, respectively, with a balanced precision-recall profile. The techniques of Explainable AI (XAI) were used to increase the transparency of this model by identifying areas in the image that are significant to the actual lesion. The framework was further designed with practical web-based deployment in mind, evaluating both lightweight and high-capacity models to balance computational efficiency against predictive accuracy. From this research, it was concluded that using model fusion, data augmentation, and segmentation-aware lesion detection would provide a solution for managing guava diseases effectively.

11
Effect of Wall Motion Sampling on CFD-Derived Left Atrial Flow Metrics

Stöcker, Y.; Guerrero-Hurtado, M.; Duran, E.; Gonzalo, A.; Ristic, Z.; Telle, A.; Kassar, A.; Haykal, R.; Akoum, N.; Boyle, P. M.; Flores, O.; Augustin, C. M.; del Alamo, J. C.; Garcia-Villalba, M.

2026-07-31 bioengineering 10.1101/2025.10.24.684343 medRxiv
Top 0.3%
7.8%
Show abstract

The temporal resolution of medical imaging sequences used to drive patient-specific computational fluid dynamics (CFD) simulations remains limited, typically providing 10-20 frames per cardiac cycle. Therefore, temporal interpolation to reconstruct left atrial (LA) wall motion and boundary conditions is required, but its impact on hemodynamic predictions has not been systematically characterized. To investigate this, we constructed high-temporal-resolution reference wall-motion data using electromechanical (EM) simulations on five patient-specific atrial geometries with a history of atrial fibrillation. We then generated temporally downsampled datasets to emulate clinical frame rates (5, 10, 20, and 40 frames per cycle) and performed CFD simulations to isolate the effects of temporal undersampling on hemodynamic metrics. The focus was placed on kinetic energy, KE, and residence time, TR, particularly in the left atrial appendage (LAA), where thrombosis is most likely to occur. We employed an immersed boundary method to prescribe the wall motion and computed blood TR through a passive scalar transport equation. Results indicate that while global LA hemodynamic indices were marginally affected by the frame rate (errors < 9%), LAA metrics were more sensitive with errors up to 31% compared to reference values. The results based on 20 and 40 frames per cycle yielded favorable agreement with reference results, while 5-and 10-frame reconstructions showed larger, though not systematically biased, deviations from the reference. Importantly, patient ranking by blood-stasis indices was largely preserved. The analysis suggests that patientspecific LA reconstructions derived from dynamic CT imaging provide a reliable basis for estimating LAA blood-stasis indices. Higher frame rates ([&ge;] 20 per cycle) offer improved quantitative accuracy, while lower temporal resolutions may remain informative for patient stratification purposes, where relative ranking is more relevant than absolute accuracy.

12
Effects of Left Atrial Wall Thickness on Myocardial Mechanics and Blood Dynamics using Multiscale Modeling

Gan, B.; Shi, L.; Chen, I. Y.; Vedula, V.

2026-06-12 bioengineering 10.64898/2026.06.09.731221 medRxiv
Top 0.3%
7.7%
Show abstract

PurposePatient-specific models of left atrial (LA) mechanics often assume uniform left atrial wall thickness (LAWT), but the effect of LAWT on the mechanics and hemodynamics remains less quantified. MethodsFour LA myocardium models were built from gated CTA images: a baseline variable thickness (VT#0), two reduced-dilation variants, and a 2mm uniform thickness model. Multi-scale mechanics and blood flow simulations were performed across all the thickness variants using model parameters personalized on the baseline model. Predicted displacements, wall stresses and strains, and hemodynamics were compared. ResultsAcross all LAWT variants, myocardial volume spanned 14.4-19.9mL (38%), while cavity volume remained mostly within 5% of image data throughout the cardiac cycle. Circulatory system output, myocardial displacements, and strains varied by 5-6% relative to the baseline model. Instantaneous stresses increased by up to 19% in the thinner variable thickness models and decreased by up to 16% in the uniformly thick case. Globally, the area under low time-averaged wall shear stress (TAWSS) varied between 23% and 30% across all thickness variants, while LA exposed to elevated oscillatory shear index (OSI) increased from nearly 6% to 19%. Over 90% of LAA was exposed to low shear, but the high-OSI area increased from 7% in VT#0 to over 30% in Uniform. ConclusionA personalized multiscale modeling framework was leveraged to demonstrate that the left atrial myocardial stresses and oscillatory shear had a greater sensitivity to local wall thickness representation compared to cavity volumes, tissue displacements, strains, and mean blood shear.

13
Foundation Model RNAGAN Enhances Biomedical Insight of Nasopharyngeal Carcinoma Metastasis

Hou, Z.; Qian, Y.; Lee, V. H.-F.; Kwong, D. L.-W.; Guan, X.; Liu, Z.; Dai, W.

2026-07-09 cancer biology 10.64898/2026.07.02.736240 medRxiv
Top 0.3%
7.7%
Show abstract

RNAGAN (version 2.0, https://github.com/ZhaozhengHou-HKU/RNAGAN-2.0.git) is a published foundation model that analyzes single-cell and bulk-level RNA sequencing samples and enables multiple applications that enhance medical insights. Here we applied this model to Nasopharyngeal Carcinoma (NPC) as in-context few-short format (i.e., the model was never trained with any NPC data). We conducted all four supported functions, which include sample stratification, vectorization, pseudo data generation, and marker identification. The results were then used for identifying metastatic NPC and to investigate mechanisms associated with NPC metastasis. Examination with stratification showed that the accuracy of RNAGAN results for evaluating the metastasis risk in NPC patients are comparable to or outcompeted recently published risk estimation linear prediction model. Vectorization results present consistency across multiple cohorts and RNAGAN model versions. In the task of identifying markers and mechanisms related to NPC metastasis, incorporating pseudo data substantially enhanced the representativeness of single-cohort-based differential expression (DE) analysis. Moreover, RNAGAN identified metastasis-related marker genes based on single cohort, were concordant with the ground truth obtained across multiple cohorts (p=1.05e-9). Regarding biomedical mechanisms, RNAGAN enabled second-order feature extraction, unveiling a remarkable domination of the protective function of adaptive immune responses (as indicated by IL21R levels) over the hazardous function of chronic, non-resolving innate inflammation (as indicated by S100A8 levels) against NPC metastasis after first-line treatment. This association demonstrates a high degree of consistency with the external cohort. This study demonstrates the utility of the foundation model RNAGAN in uncovering therapeutic insights for novel cancer types without extra training. We reveal a critical spatial mechanism preventing distant metastasis via humoral anti-tumor immunity in NPC. High S100A8 expression by innate antigen-presenting cells (APCs) triggers an inflammatory cascade promoting epithelial-mesenchymal transition (EMT) and metastasis. However, when germinal center IL21R+ B cells simultaneously colocalize with these innate signals, they override this suppressive tissue stress. Spatial analysis shows that a high S100A8/IL21R intersection within tumor regions strictly distinguishes treatment responders, whereas non-responders display spatial mismatch or S100A8+ hyper-infiltration. This coordinated innate-adaptive cross-talk sustains functional tertiary lymphoid structures (TLS) that mature IgG-secreting plasma cells, which opsonize and eliminate emerging EMT tumor cells before systemic escape. Consequently, while S100A8 alone is an unreliable prognosticator, its spatial colocalization with IL21R is a robust protective indicator overlooked by conventional bulk analysis methods.

14
Computational Fluid Particle Dynamics (CFPD)-Based Virtual Next Generation Impactor (vNGI) to Predict the Aerodynamic Particle Size Distribution (APSD) of Respiratory Drug Delivery Products: Toward New Approach Methodologies (NAMs) in Inhaler Performance Evaluation

Patil, A. S.; Feng, Y.

2026-06-30 bioengineering 10.64898/2026.06.29.735263 medRxiv
Top 0.3%
7.6%
Show abstract

The Next Generation Impactor (NGI) is one of the regulatory gold standards for characterizing aerodynamic particle size distributions (APSDs) of orally inhaled drug products (OIDPs); however, its reliance on complex, resource-intensive in vitro testing under tightly controlled environmental conditions limits experimental flexibility and introduces variability. In alignment with the growing regulatory emphasis on New Approach Methodologies (NAMs) for drug development, this study presents a rigorously validated computational fluid particle dynamics (CFPD) based virtual NGI (vNGI) as an in silico method complementary to conventional testing. The vNGI replicates a significant portion of the NGI geometry and airflow physics, enabling high-resolution spatiotemporal analysis of aerosol transport and deposition mechanisms that are otherwise inaccessible experimentally. A comprehensive verification and validation framework was implemented, including mesh and particle independence studies, turbulence model assessment, and comparison of stagewise deposition efficiencies with available in vitro data at 30 L/min. The model's capabilities were further extended to low and high flow rates, and two bio-relevant mouth-throat models and polydisperse particle laden aerosol were added. The model demonstrates strong predictive capability for a few stages and provides mechanistic insight into discrepancies in other stages, depending on the type of analysis. Importantly, this work establishes the vNGI as a fit-for-purpose according to NAM by (i) defining a clear context of use for APSD prediction and inhaler performance evaluation, (ii) capturing physically and biologically relevant air-particle interactions, and (iii) demonstrating technical robustness and reproducibility through systematic validation. The platform can potentially further enable simulation of environmental and physiological conditions, such as humidity effects, that are difficult to control experimentally, thereby improving human relevance and reducing reliance on costly and time-consuming in vitro testing. This study positions the vNGI as a scalable, regulatory aligned NAM capable of supporting early stage drug device combination product development, device optimization, and an alternative bioequivalence assessment, contributing to ongoing efforts to enhance predictive performance, reduce experimental burden, and transition toward human centric, inhalation product evaluation.

15
Revisiting Logistic Regression for High-Dimensional Gene Expression Data

Souza, R. d. O.; Rodrigues, W. F.; Couto, B.; Dos Santos, M. A.

2026-07-24 bioinformatics 10.64898/2026.07.20.739668 medRxiv
Top 0.3%
7.3%
Show abstract

Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.

16
Personalized planning of cardiac resynchronization therapy through integration of coronary sinus geometry, clinical data, digital twins, and machine learning: visualization, stratification, and optimization

Bazhutina, A.; Chumarnaya, T.; Zubarev, S.; Budanova, M.; Stepanova, V.; Khamzin, S.; Lebedev, D.; Solovyova, O.

2026-07-02 cardiovascular medicine 10.64898/2026.07.01.26356827 medRxiv
Top 0.3%
7.3%
Show abstract

Background: Cardiac resynchronization therapy (CRT) fails in 30% of patients, often due to suboptimal left ventricular pacing site (LVPS) selection. Current practice lacks tools for pre-procedural, patient-specific LVPS optimization within the accessible coronary sinus (CS) tributaries. This study aimed to develop a digital twin and an explainable ML-based clinical decision support framework to address this issue. Methods: Personalized 3D cardiac models incorporating ventricular anatomy, myocardial fibrosis, and CS anatomy were constructed from CT and LGE-MRI for 74 CRT candidates. Finite-element Eikonal simulations of biventricular pacing generated patient-specific electrophysiological features at candidate LVPS. A Machine Learning (ML) classifier was trained on a hybrid feature set of pre-procedural clinical variables and model-derived indices, validated by leave-one-out cross-validation. SHAP analysis provided a physiologically interpretable rationale for each prediction. The framework was applied to a pilot cohort of 19 patients with reconstructed 3D CS anatomy to generate a spatial likelihood map of CRT response across all clinically implantable pacing sites within each patient's CS. Results: The ML classifier outperformed the reference Feeny clinical calculator under LOO-CV (accuracy 0.78 vs 0.58; F1-score 0.75 vs 0.43), AUC=0.78, sensitivity=0.80, specificity=0.77. Bootstrap analysis yielded mean AUC=0.85 (95% CI 0.70-0.95). In the pilot CS cohort, the framework identified that 8 of 13 clinical non-responders had no accessible CS site predicted to yield a positive response, supporting redirection towards alternative pacing strategies. In the remaining 5, alternative implantable sites with high predicted response probability were identified. SHAP analysis confirmed that dominant predictors were patient-specific in their relative contributions, supporting individualized over heuristic-based LVPS selection. Conclusion: This pilot study demonstrates the feasibility of a digital twin and explainable ML framework as a pre-procedural clinical decision support tool for CRT planning, stratifying patients and identifying optimal implantable sites with transparent anatomical rationale. Prospective validation and regulatory evaluation are required before clinical deployment.

17
A Deep Hypergraph Learning Model for Predicting Antimicrobial Combination Effects Across Bacterial Targets

Midjani, F.; Rajabi, A. H.; Keshtkar, F. Z.; Malekpour, M.; Jafarizadeh, A.; Alizadehsani, R.; Plawiak, P.

2026-06-11 bioinformatics 10.64898/2026.06.09.731104 medRxiv
Top 0.3%
7.2%
Show abstract

Antimicrobial resistance (AMR) creates an urgent need for efficient strategies to identify effective antibacterial combinations. Combination therapy, including antimicrobial peptides (AMPs) paired with conventional antibiotics, is a promising approach, but exhaustive experimental screening across drug pairs and bacterial targets is impractical. This study introduces a hybrid GCN-based hypergraph neural network (HGNN) for predicting antimicrobial-agent combination outcomes against bacterial targets. Each antimicrobial-agent-antimicrobial-agent-bacterium triplet is represented as a ternary hyperedge, enabling the model to learn context-dependent interaction patterns. The framework integrates SMILES-derived molecular graph embeddings for antimicrobial agents, including conventional antibiotics and AMPs, with taxonomy-derived bacterial representations. The prediction task was formulated as a three-class classification problem: synergy, antagonism, and non-interaction. The non-interaction class included experimentally verified indifferent records and synthetic presumed non-interaction triplets generated by negative sampling. Model development used drug-pair-grouped splitting, five-fold grouped cross-validation within the training/validation partition, and final evaluation on a held-out test set. On the held-out three-class test set, the selected GCN-based HGNN achieved an accuracy of 0.83, weighted F1-score of 0.84, macro F1-score of 0.80, and ROC-AUC of 0.95. Per-class evaluation showed accuracies of 0.80 for synergy, 0.92 for antagonism, and 0.85 for non-interaction. Pair-type analysis showed strong performance across AMP-AMP, AMP-conventional antibiotic, and conventional antibiotic-conventional antibiotic combinations. These findings suggest that hypergraph-based representation learning can support computational prioritization of antimicrobial combinations for experimental follow-up. Further studies will be needed to improve model interpretability and to perform prospective validation of predicted synergistic combinations.

18
Artificial Intelligence Model: Optimizing Cancer Risk Level Predictions Using Machine learning and deep learning approaches

Abd Aziz, A. B.; Arabiat, A.; Abu Owida, H.; Abuowaida, S.; Alshdaifa, N.; A. Mashagba, H.

2026-08-25 cancer biology 10.64898/2026.08.20.745910 medRxiv
Top 0.3%
7.0%
Show abstract

This study emphasizes the potential of computational techniques in cancer risk assessment, lighting opportunities for specific and data-driven healthcare solutions. This study examines the use of artificial intelligence (AI), machine learning (ML), and deep learning (DL) approaches to improve cancer risk assessment using a Kaggle dataset. The study uses Java-based ML software to create and evaluate multiple predictive models, taking advantage of its powerful libraries and frameworks for processing and analyzing cancer risk indicators. This work analyzes model performance using 10-fold cross-validation, resulting in reliable generalization and accuracy estimates. Several classification techniques, such as Random Forest (RF) logistic regression (LR), decision trees (DT), Naive Bayes (NB), and Multi-layer perceptron (MLP), are used to assess their efficacy in predicting risk levels for various cancer types. To measure classification effectiveness, key performance metrics such as accuracy, precision, recall, and F1 score are produced, in addition to multi-class confusion matrices. The results show that the RF model is the best classifier for classification, with accuracy of 99.85%, F-measure of 99.80%, precision of 99.80%, and sensitivity of 99.90%. These findings demonstrate the model's ability to effectively estimate cancer risk levels among individuals. of cancer risk estimations, allowing for earlier discovery and more effective medical care.

19
Data-Driven Stochastic Model for Detecting Patientswith Alzheimer's Disease

Abeywardana, G. D.; Tsokos, C.

2026-06-15 neurology 10.64898/2026.06.06.26355081 medRxiv
Top 0.3%
6.9%
Show abstract

Alzheimer s disease (AD) is a critical neurological disorder that causes the brain to shrink and leads to the eventual death of brain cells, adversely affecting a person s ability to function. AD is a fast-growing disease in the United States and was the fifth leading cause of death among Americans 65 years of age or older in 2023. In the United States 6.9 million people aged 65 or older were diagnosed with AD, along with a high rate of undiagnosed patients. Thus, the objective of our study is to develop a real data-driven predictive model to identify a patient with AD based on eight risk factors: Age, Gender, ADAS-Cog13, Entorhinal, Fusiform, Intracranial Volume (ICV), Amyloid-Beta, and Tau Protein, with a high degree of accuracy. The quality of the model was evaluated using well-established and sophisticated statistical measures: the area under the receiver operating characteristic curve, calibration plot, Hosmer-Lemeshow goodness-of-fit test, and K-fold cross-validation. If a patient is given information on the above risk factors, our proposed binary logistic regression model can classify the patient as having AD or not with at least 98% accuracy.

20
Consistency Analyses of Open-source Software for Motor Unit Decomposition Using High-density Electromyography Signal

Fu, J.; Zhang, S.; Huang, H. J.; Rakhshan, M.; Wen, Y.

2026-07-08 bioengineering 10.64898/2026.07.07.737019 medRxiv
Top 0.4%
6.8%
Show abstract

Motor unit (MU) decomposition using high-density surface electromyography (HD-sEMG) has been widely used to characterize MU behavior in neurophysiology and to build neural-machine interfaces for wearable robots. Recently, many open-source software tools for MU decomposition have been made available on GitHub, which could reduce the effort of researchers in the field. However, the consistency among these open-source tools has never been studied, making researchers hesitate to use them. In this study, we collected 7 open-source software tools on GitHub and applied them to decompose MUs from an open-source HD-sEMG dataset (including 11 isometric contraction trials) to investigate the consistency among these tools. To create a comprehensive MU pool for reference, we combined all unique MUs identified by seven tools, visually inspected and removed bad MUs, and manually edited all remaining MU spike trains. Across 7 tools for 11 trials, the number of identified MUs ranges from 167 to 736. The number of valid MUs after expert inspection ranges from 29 to 210, which is 10% to 72% of the reference pool. The rate of agreement between the raw MUSTs and the manually edited MUSTs ranges from 0.86 to 0.94, and the averaged number of edits per MU to correct misalignments ranges from 14 to 39. The results show inconsistency in the implementation and procedures of each tool, which results in an inconsistent number of identified MUs and valid MUs (29 vs 210). In general, a substantial amount of effort is required to process the raw MUSTs from each tool to conduct further research analysis. This study provided a guideline for using open-source software tools for MU decomposition and indicated that it would be beneficial to develop tools to automatically edit the MUSTs.